Segmenting Arabic Handwritten Documents into Text lines and Words
نویسندگان
چکیده
In this paper, we present a method for segmenting Arabic handwritten documents into text lines and words. Text line segmentation is addressed by a well-known technique, the horizontal projection profile, in which autocorrelation is used to enhance the self similarity of this profile. This technique promotes the estimation of text line spacing. Word extraction is based on an adaptation of a known method, gap metrics.This improvement relies on deriving the values of these gaps from the properties of each input document, making the proposed method tolerant and robust to Arabic handwritten nature. Text is often divided into words, sub-words and letters; however, some letters do not connect to the following letter, even in the middle of a word. A gap metric method exploits the membership values of a clustering algorithm to identify segmentation thresholds as “within word” or “between words” gaps. The proposed method is tested on the benchmarking datasets of Arabic handwritten text recognition research (AHDB), and very promising results were achieved, with an 84.8% correct extraction rate.
منابع مشابه
Word Extraction and Recognition in Arabic Handwritten Text
Segmenting arabic manuscripts into text-lines and words is an important step to make recognition systems more efficient and accurate. The major problem making this task crucial is the word extraction process: first, words are often a succession of sub-words where the space value between these sub-words do not respect any rules. Second, the presence of connections even between non adjacent sub-w...
متن کاملComponent-based Segmentation of Words from Handwritten Arabic Text
Efficient preprocessing is very essential for automatic recognition of handwritten documents. In this paper, techniques on segmenting words in handwritten Arabic text are presented. Firstly, connected components (ccs) are extracted, and distances among different components are analyzed. The statistical distribution of this distance is then obtained to determine an optimal threshold for words se...
متن کاملSegmentation of Handwritten and Printed Arabic Documents
on this paper, we proposed a new text line segmentation of handwritten and typewriting Arabic document images that uses the Outer Isothetic Cover (OIC) algorithm of a digital object. In the first step, we use this method to segment the composed document into text blocs. In the second step, for each text bloc we will extract the text lines. Finally, line text will be segmented into words or into...
متن کاملHolistic Approach for Classifying and Retrieving Personal Arabic Handwritten Documents
This paper presents a novel holistic technique for classifying and retrieving Arabic handwritten text documents. The retrieval of Arabic handwritten documents is performed in several steps. First, the Arabic handwritten document images are segmented into words, and then each word is segmented into its connected parts. Second, several features are extracted from these connected parts and then co...
متن کاملRegion growing based segmentation algorithm for typewritten and handwritten text recognition
This paper presents a new technique of high accuracy to recognize both typewritten and handwritten English and Arabic texts without thinning. After segmenting the text into lines (horizontal segmentation) and the lines into words, it separates the word into its letters. Separating a text line (row) into words and a word into letters is performed by using the region growing technique (implicit s...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2014